Papers by Arun Balaji Buduru

4 papers
FAtNet: Cost-Effective Approach Towards Mitigating the Linguistic Bias in Speaker Verification Systems (2022.findings-naacl)

Copied to clipboard

Challenge: Linguistic bias in Deep Neural Network (DNN) based systems is a critical challenge that needs attention.
Approach: They propose to integrate a lightweight embedding with existing NLP systems to mitigate linguistic bias without adaptation.
Outcome: The proposed framework reduces linguistic bias and enhances usability of baselines for twelve languages.
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages (2026.findings-acl)

Copied to clipboard

Challenge: Speech deepfakes are highly realistic and can generate a few seconds of recorded speech.
Approach: They propose an ALM that integrates semantic and prosodic representations from Whisper and TRILLsson to generate a speech deepfake dataset.
Outcome: The proposed framework outperforms existing ALMs on the ICF benchmark in Indic languages.
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution (2025.findings-acl)

Copied to clipboard

Challenge: x-vector (speaker recognition PTM) achieves the highest performance in prosodic tasks . despite its low parameter, x vector captures unique prosodic characteristics of the sources .
Approach: They propose to use SOTA speech pre-trained models to capture prosodic sig-natures of generative sources for audio deepfake source attribution.
Outcome: The proposed model captures prosodic sig-natures of generative sources better than other models on ASVSpoof and CFAD.
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)

Copied to clipboard

Challenge: a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information.
Approach: They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection .
Outcome: The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations